Back

International Journal of Epidemiology

Oxford University Press (OUP)

All preprints, ranked by how well they match International Journal of Epidemiology's content profile, based on 88 papers previously published here. The average preprint has a 0.06% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Mendelian randomization infers the effect of 14 parental illnesses on 44 congenital anomalies

Li, Y.; Shao, W.; Tian, T.; Tan, L.

2024-07-14 sexual and reproductive health 10.1101/2024.07.13.24310358 medRxiv
Top 0.1%
62.6%
Show abstract

BackgroundCongenital anomalies (CA), including congenital malformations (CM) and congenital deformities (CD), are significant health concerns influenced by genetic and environmental factors. Parental illnesses, especially those with genetic components, may affect the risk of congenital anomalies in offspring. Although clinical studies have suggested associations between certain parental illnesses and increased CM and CD risk, causal relationships remain unclear. This study employs a Mendelian randomization (MR) approach to investigate these potential causal links. MethodsFourteen parental illnesses were selected for this study: breast cancer, chronic bronchitis/emphysema, diabetes, heart disease, hypertension, and Alzheimers disease in mothers; and Alzheimers disease, bowel cancer, chronic bronchitis/emphysema, diabetes, heart disease, hypertension, lung cancer, and prostate cancer in fathers. Genetic variants associated with these illnesses were identified from genome-wide association studies (GWAS) in the UK Biobank. Genetic data for 44 congenital anomalies were sourced from the FinnGen database. Two-sample MR was conducted to estimate causal effects, with sensitivity analyses and multivariable MR (MVMR) to control for potential confounders. ResultsMR analysis revealed causal relationships between 13 parental illnesses and 13 specific congenital anomalies. Notably, mothers hypertension significantly increased the risk of congenital hypothyroidism (IVW: OR = 7.969, 95% CI = 3.0826-20.6011, p = 4.20E-04), and fathers diabetes increased the risk of congenital heart defects in offspring (IVW: OR = 3.8E+09, 95% CI = 2.2E+04-6.6E+14, p = 3E-04). The associations strength varied with the type of parental illness and the specific congenital disease. ConclusionThis study underscores the utility of MR in elucidating genetic influences of parental health conditions on congenital anomalies. The findings highlight the importance of managing parental health to reduce congenital anomalies risk in offspring. Further research is needed to explore underlying biological mechanisms and validate these findings in diverse populations.

2
Investigating the potential causal relationship between parity and long-term maternal cardiometabolic health outcomes using Mendelian randomization

Brito Nunes, C.; Fraser, A.; Moen, G.-H.; Hatton, A. A.; Evans, D.

2026-08-11 sexual and reproductive health 10.64898/2026.08.09.26360053 medRxiv
Top 0.1%
49.9%
Show abstract

Background: Multiple observational studies have reported associations between greater parity and increased CVD risk. Whether these associations reflect causal effects or are confounded by socioeconomic factors remains unclear. Methods: We investigated associations between number of children ever born (NEB) and 16 cardiometabolic traits in up to 172,122 females and 138,390 males in the UK Biobank, and an independent sample of 53,237 UK Biobank spousal pairs. We additionally conducted sex-stratified two-sample Mendelian randomization (MR) and applied a novel spousal MR framework, in which an individual's spouse's genotype was used as the instrumental variable to estimate the causal effect of NEB on cardiometabolic health outcomes, as an approach to minimize bias from horizontal pleiotropy. Results: NEB was associated with multiple cardiometabolic traits in the multivariable regression, even after adjustment for socioeconomic status, with differences in the strength of association observed between males and females. Traditional MR provided evidence that higher NEB causally increases type 2 diabetes risk in females, body mass index (BMI) in both sexes, female basal metabolic rate (BMR) and male body fat percentage but decreases female blood pressure. Spousal MR corroborated positive effects on female BMI and BMR and additionally suggested inverse causal effects on female HDL cholesterol and ApoA1 and male blood glucose. Conclusion: These findings indicate a possible causal relationship between NEB and long-term cardiometabolic health, although causal effects are likely to be small.

3
From Genomics Data to Causality: An Integrated Pipeline for Mendelian Randomization

Sharma, J.; Jangale, V.; Swain, A. K.; Yadav, P.

2023-11-04 epidemiology 10.1101/2023.11.04.23298053 medRxiv
Top 0.1%
39.5%
Show abstract

BackgroundMendelian randomization (MR) has emerged as a valuable tool for causal inference in genetic epidemiology. Existing MR methods have issues related to pleiotropy and offer limited comprehensiveness. Here, we introduce an integrated MR analysis pipeline designed for GWAS summary statistics data. Our pipeline integrates feature selection, harmonization, and checkpoint mechanisms to improve the accuracy and reliability of MR analysis. MethodsIn classical GWAS, the p-value threshold usually does not guarantee to identify causal single-nucleotide polymorphisms (SNPs). In such cases, t-statistics can be considered as imperative and robust criteria for identifying causal SNPs. Therefore, in this study, we computed the t-statistic for all independent SNPs remained after linkage disequilibrium pruning. Next, prior to harmonization, we removed SNPs having a t-statistic below the average t-statistic value. Furthermore, our pipeline incorporates sensitivity analysis tests at each step to reduce the chances of directional pleiotropy. Result and ConclusionWe applied our pipeline to single-sample and two-sample MR study designs, encompassing diverse populations and a wide range of diseases. Our results demonstrate superior performance compared to existing MR methods. In conclusion, our research presents an integrated MR analysis pipeline that significantly enhances the accuracy and reliability of MR studies. By outperforming existing methods and providing comprehensive validation, this pipeline represents a valuable tool for researchers in genetics and epidemiology.

4
Re-evaluating the robustness of Mendelian randomisation to measurement error

Woolf, B.; Karhunen, V.; Yarmolinsky, J.; Tilling, K.; Gill, D.

2022-10-04 epidemiology 10.1101/2022.10.02.22280617 medRxiv
Top 0.1%
38.6%
Show abstract

BackgroundMendelian randomisation (MR) uses germline genetic variation as a natural experiment to investigate causal relations between traits. MR is robust to non-differential random measurement error in exposures or outcomes. However, the effect of differential measurement error, and non-differential measurement error on the variant selection process, remains unclear. MethodsWe use Monte-Carlo simulations and an applied example to explore the effect of differential measurement error on MR estimates for a continuous exposure and outcome, and the application of multivariable MR to reduce bias. We then explore the effect of non-differential measurement error during variant selection on MR analysis, using simulated and real-world data in the UK Biobank. ResultsCausal differential measurement error biased MR estimates when it occurred in the outcome, or in an exposure with a true causal effect on the outcome. This bias was mitigated by including the variable causing the error in a multivariable MR analysis. Unlike standard regression, MR was not biased by non-causal differential measurement error, i.e. when a third variable caused the exposure (or outcome) and the error in the outcome (or exposure). Non-differential measurement error in the phenotype during variant selection reduced the precision of MR estimates and induced bias. This bias was attenuated by using three-sample MR, or Winners curse corrections. ConclusionMR estimates can be biased by differential measurement error, but in fewer circumstances than standard regression. Multivariable MR can be used to attenuate differential measurement error if the error mechanism is known. Three-sample MR is recommended particularly for error-prone exposures. Key MessagesO_LIPrevious research demonstrates that Mendelian randomization (MR) is unbiased by (classical) non-differential measurement error in the exposure or outcome once the genetic instruments have been identified. C_LIO_LIMR estimates can be biased by causal differential measurement error in a continuous outcome, or in a continuous exposure when there is a true causal effect of the exposure on the outcome. As with observational studies, this bias could lead to an over-or under-estimation of the true effect estimate. C_LIO_LIUnlike standard regression, MR is not biased by non-causal differential measurement error between the exposure and outcome, or causal differential measurement error in the exposure under the null hypothesis. C_LIO_LIWhen all the requisite assumptions are met, multivariable MR can be used to attenuate bias due to differential measurement error in an exposure or outcome, if the variables causing the error are known. Else, a smaller sample, which is less susceptible to differential measurement error, would produce more accurate estimates, despite decreased power. C_LIO_LINon-differential measurement error in the exposure will reduce precision and can cause bias in MR when it occurs during the instrument selection process. The bias caused by non-differential measurement error in instrument selection can be mitigated by using non-overlapping samples for instrument selection and the instrument-exposure estimation, or statistical correction for Winners curse. C_LI

5
Survival bias and competing risk can severely bias Mendelian Randomization studies of specific conditions

Schooling, C. M.; Lopez, P. M.; Au Yeung, S.; Hunag, J. V.

2019-07-26 genetics 10.1101/716621 medRxiv
Top 0.1%
35.3%
Show abstract

BackgroundMendelian randomization (MR) provides unconfounded estimates. MR is open to selection bias particularly when the underlying sample is selected on surviving the genetically instrumented exposure and other conditions that share etiology with the outcome (competing risk before recruitment). Few methods to address this bias exist. MethodsWe use directed acyclic graphs to show this selection bias can be addressed by adjusting for common causes of survival and outcome. We use multivariable MR to obtain a corrected MR estimate, specifically, the effect of statin use on ischemic stroke, because statins affect survival and stroke typically occurs later in life than ischemic heart disease so is open to competing risk. ResultsIn univariable MR the genetically instrumented effect of statin use on ischemic stroke was in a harmful direction in MEGASTROKE and the UK Biobank (odds ratio (OR) 1.33, 95% confidence interval (CI) 0.80 to 2.20). In multivariable MR adjusted for major causes of survival and ischemic stroke, (blood pressure, body mass index and smoking initiation) the effect of statin use on stroke in the UK Biobank was as expected (OR 0.81, 95% CI 0.68 to 0.98) with a Q-statistic indicating absence of genetic pleiotropy or selection bias, but not in MEGASTROKE. ConclusionMR studies concerning late onset chronic conditions with shared etiology based on samples recruited in later life need to be conceptualized within a mechanistic understanding, so as to any identify potential bias due to competing risk before recruitment, and to inform the analysis and interpretation.

6
Age-varying genetic associations and implications for bias in Mendelian randomization analyses

Labrecque, J. A.; Swanson, S. A.

2021-04-30 epidemiology 10.1101/2021.04.28.21256235 medRxiv
Top 0.1%
35.2%
Show abstract

Estimates from conventional Mendelian randomization (MR) analyses can be biased when the genetic variants proposed as instruments vary over age in their relationship with the exposure. For four exposures commonly studied using MR, we assessed the degree to which their relationship with genetic variants commonly used as instruments varies by age using flexible, spline-based models in UK Biobank data. Using these models, we then estimated how biased MR estimates would be due to age-varying relationships using plasmode simulations. We found that most genetic variants had age-varying relationships with the exposure for which they are a proposed instrument. Body mass index and LDL cholesterol had the most variation while alcohol consumption had very little. This variation over age led to small potential biases in some cases (e.g. alcohol consumption and C-reactive protein) and large potential biases for many proposed instruments for BMI and LDL.

7
Data Resource Profile: Genomic Data in Multiple British Birth Cohorts (1946-2001) - Health, Social, and Environmental Data from Birth to Old Age

Shireby, G.; Morris, T. T.; Wong, A.; Chaturvedi, N.; Ploubidis, G. B.; Fitzsimmons, E.; Goodman, A.; Sanchez-Galvez, A.; Davies, N. M.; Wright, L.; Bann, D.

2024-11-06 epidemiology 10.1101/2024.11.06.24316761 medRxiv
Top 0.1%
35.0%
Show abstract

Birth cohort studies have a rich history of contributing to science across disciplinary fields, notably health and social sciences. Here, we introduce a curated resource comprising genomic data from five British birth cohort studies--longitudinal studies with extensive data collected prospectively across life, each deliberately sampled to be nationally representative (born 1946-2001). These contain health and social data from birth to older age, enabling longitudinal and cross-cohort genetically informed research. The Millennium Cohort Study additionally includes data on parents and offspring, enabling within-family analyses. Across five cohorts born in 1946, 1958, 1970, 1989-90, and 2000-2002, 27,432 participants have harmonized, imputed, and quality-controlled genetic data from genotyping arrays covering 6.7 million common SNPs. The Millennium Cohort Study contains over 6,000 mother-offspring pairs and over 3,000 mother-father-offspring trios. Pseudonymized data are freely available to the global research community upon approval of a data access request (https://cls.ucl.ac.uk/data-access-training).

8
Cohort Profile Update: Expanding the Cardiovascular Risk in Young Finns Study into a multigenerational cohort

Pahkala, K.; Rovio, S.; Auranen, K.; Bourgery, M.; Elovainio, M.; Fogelholm, M.; Haapala, J.; Hirvensalo, M.; Hutri, N.; Jokinen, E.; Jula, A. M.; Juonala, M.; Kaikkonen, J.; Kartiosuo, N.; Kiviranta, H.; Koskinen, J. S.; Kotaja, N.; Kahonen, M.; Laitinen, T. P.; Lehtimaki, T.; Lisinen, I.; Loo, B.-M.; Lyytikainen, L.-P.; Magnussen, C. G.; Mishra, P. P.; Mykkanen, J.; Makela, J.-A.; Mannisto, S.; Nevalainen, J.; Pulkki-Raback, L.; Raitoharju, E.; Rantakokko, P.; Ronnemaa, T.; Stenbacka, S.; Taittonen, L.; Tammelin, T. H.; Toppari, J.; Tossavainen, P.; Viikari, J.; Raitakari, O.

2024-11-26 cardiovascular medicine 10.1101/2024.11.24.24317769 medRxiv
Top 0.1%
34.9%
Show abstract

Cardiovascular Risk in Young Finns Study (YFS) is a prospective cohort of 3596 males and females (baseline age 3-18 years) that was established in 1980 to study cardiovascular risk factors in children and adolescents. The YFS has been instrumental in demonstrating the links between childhood risk factors and adult cardiovascular outcomes with implications on paediatric cardiovascular preventive practise. In the latest follow-up in 2018-2020, the study was expanded into a three-generation cohort, including the original cohort members, as well as their parents and offspring. Altogether 7341 individuals aged 3-92 years participated; 2127 original cohort members (now aged 40-58 years), 2452 parents (aged 58-92 years) and 2762 offspring (aged 3-37 years) of the original cohort members. The main aim was to establish a multigenerational population data base, sample biobank and links to national health registries that would offer a unique platform to study familial transmission of health-related traits. The field studies and data collections were conducted as part of the ERC funded MULTIEPIGEN project designed testing a specific hypothesis that epigenetic markers in the semen play a role in the transmission of intergenerational information in the paternal lineage. It is a first a priori designed epidemiologic study assessing the role of paternal life exposures, including chemical and psychosocial stress, in the development of cardio-metabolic, cognitive and psychosocial outcomes in their offspring.

9
The relationship between BMI and COVID-19: exploring misclassification and selection bias in a two-sample Mendelian randomisation study

Clayton, G. L.; Goncalves Soares, A.; Goulding, N.; Borges, M. C. G.; Holmes, M.; Davey Smith, G.; Tilling, K.; Lawlor, D. A.; Carter, A. R.

2022-03-05 epidemiology 10.1101/2022.03.03.22271836 medRxiv
Top 0.1%
34.8%
Show abstract

ObjectiveTo use the example of the effect of body mass index (BMI) on COVID-19 susceptibility and severity to illustrate methods to explore potential selection and misclassification bias in Mendelian randomisation (MR) of COVID-19 determinants. DesignTwo-sample MR analysis. SettingSummary statistics from the Genetic Investigation of ANthropometric Traits (GIANT) and COVID-19 Host Genetics Initiative (HGI) consortia. Participants681,275 participants in GIANT and more than 2.5 million people from the COVID-19 HGI consortia. ExposureGenetically instrumented BMI. Main outcome measuresSeven case/control definitions for SARS-CoV-2 infection and COVID-19 severity: very severe respiratory confirmed COVID-19 vs not hospitalised COVID-19 (A1) and vs population (those who were never tested, tested negative or had unknown testing status (A2)); hospitalised COVID-19 vs not hospitalised COVID-19 (B1) and vs population (B2); COVID-19 vs lab/self-reported negative (C1) and vs population (C2); and predicted COVID-19 from self-reported symptoms vs predicted or self-reported non-COVID-19 (D1). ResultsWith the exception of A1 comparison, genetically higher BMI was associated with higher odds of COVID-19 in all comparison groups, with odds ratios (OR) ranging from 1.11 (95%CI: 0.94, 1.32) for D1 to 1.57 (95%CI: 1.57 (1.39, 1.78) for A2. As a method to assess selection bias, we found no strong evidence of an effect of COVID-19 on BMI in a no-relevance analysis, in which COVID-19 was considered the exposure, although measured after BMI. We found evidence of genetic correlation between COVID-19 outcomes and potential predictors of selection determined a priori (smoking, education, and income), which could either indicate selection bias or a causal pathway to infection. Results from multivariable MR adjusting for these predictors of selection yielded similar results to the main analysis, suggesting the latter. ConclusionsWe have proposed a set of analyses for exploring potential selection and misclassification bias in MR studies of risk factors for SARS-CoV-2 infection and COVID-19 and demonstrated this with an illustrative example. Although selection by socioeconomic position and arelated traits is present, MR results are not substantially affected by selection/misclassification bias in our example. We recommend the methods we demonstrate, and provide detailed analytic code for their use, are used in MR studies assessing risk factors for COVID-19, and other MR studies where such biases are likely in the available data. SummaryO_ST_ABSWhat is already known on this topicC_ST_ABS- Mendelian randomisation (MR) studies have been conducted to investigate the potential causal relationship between body mass index (BMI) and COVID-19 susceptibility and severity. - There are several sources of selection (e.g. when only subgroups with specific characteristics are tested or respond to study questionnaires) and misclassification (e.g. those not tested are assumed not to have COVID-19) that could bias MR studies of risk factors for COVID-19. - Previous MR studies have not explored how selection and misclassification bias in the underlying genome-wide association studies could bias MR results. What this study adds- Using the most recent release of the COVID-19 Host Genetics Initiative data (with data up to June 2021), we demonstrate a potential causal effect of BMI on susceptibility to detected SARS-CoV-2 infection and on severe COVID-19 disease, and that these results are unlikely to be substantially biased due to selection and misclassification. - This conclusion is based on no evidence of an effect of COVID-19 on BMI (a no-relevance control study, as BMI was measured before the COVID-19 pandemic) and finding genetic correlation between predictors of selection (e.g. socioeconomic position) and COVID-19 for which multivariable MR supported a role in causing susceptibility to infection. - We recommend studies use the set of analyses demonstrated here in future MR studies of COVID-19 risk factors, or other examples where selection bias is likely.

10
Polynomial Mendelian Randomization reveals widespread non-linear causal effects in the UK Biobank

Sulc, J.; Sjaarda, J.; Kutalik, Z.

2021-12-10 genetics 10.1101/2021.12.08.471751 medRxiv
Top 0.1%
33.7%
Show abstract

Causal inference is a critical step in improving our understanding of biological processes and Mendelian randomisation (MR) has emerged as one of the foremost methods to efficiently interrogate diverse hypotheses using large-scale, observational data from biobanks. Although many extensions have been developed to address the three core assumptions of MR-based causal inference (relevance, exclusion restriction, and exchangeability), most approaches implicitly assume that any putative causal effect is linear. Here we propose PolyMR, an MR-based method which provides a polynomial approximation of an (arbitrary) causal function between an exposure and an outcome. We show that this method provides accurate inference of the shape and magnitude of causal functions with greater accuracy than existing methods. We applied this method to data from the UK Biobank, testing for effects between anthropometric traits and continuous health-related phenotypes and found most of these (84%) to have causal effects which deviate significantly from linear. These deviations ranged from slight attenuation at the extremes of the exposure distribution, to large changes in the magnitude of the effect across the range of the exposure (e.g. a 1 kg/m2 change in BMI having stronger effects on glucose levels if the initial BMI was higher), to non-monotonic causal relationships (e.g. the effects of BMI on cholesterol forming an inverted U shape). Finally, we show that the linearity assumption of the causal effect may lead to the misinterpretation of health risks at the individual level or heterogeneous effect estimates when using cohorts with differing average exposure levels.

11
Estimating the effect of circulating vitamin D on body mass index: a Mendelian randomization study

Chadha, M.; Bell, J.; Sanderson, E.

2023-08-04 epidemiology 10.1101/2023.08.01.23293487 medRxiv
Top 0.1%
31.1%
Show abstract

BackgroundNumerous observational studies have shown an association between higher circulating 25 hydroxyvitamin D (vitamin D) and lower body mass index (BMI). Whether this represents a causal effect remains unclear. Mendelian randomization (MR) is an approach to causal inference that uses genetic variants as instrumental variables to estimate the effect of exposures on outcomes of interest. MR estimates are not biased by confounding, reverse causation and other biases in the same way as conventional observational estimates. In this study, we used MR with new data on genetic variants associated with vitamin D to estimate the effect of vitamin D on BMI. MethodsWe selected single nucleotide polymorphisms (SNPs) which were associated with vitamin D in a recent large genome-wide association study (GWAS) at genome-wide significance as instruments for vitamin D. We used inverse variance weighted models and further assessed individual SNPs that showed evidence of an effect, and biologically informed SNPs located in genetic regions previously associated with vitamin D, for associations with other traits at genome-wide significance, using Wald ratio estimation. ResultOur main results showed no evidence of an effect of vitamin D on BMI (estimated standard deviation change in BMI per standard deviation change in vitamin D: -0.003, 95% confidence interval [-0.06, 0.06]). This was also supported by pleiotropy robust sensitivity analyses. Individual SNPs that showed evidence of an effect of vitamin D on either lower or higher BMI were strongly associated with numerous other traits suggesting high levels of horizontal pleiotropy. Biologically informed SNPs showed no evidence of a causal effect of vitamin D on BMI and showed substantially less evidence of pleiotropic effects. ConclusionThe observed association between vitamin D and BMI is unlikely to be due to a causal effect of vitamin D on BMI. We also show how additional evidence can be incorporated into an MR study to interrogate individual SNPs for potential pleiotropy and improve interpretation of results.

12
Estimating and visualising multivariable Mendelian randomization analyses within a radial framework

Spiller, W.; Bowden, J.; Sanderson, E.

2023-04-04 epidemiology 10.1101/2023.04.04.23288134 medRxiv
Top 0.1%
31.0%
Show abstract

BackgroundMultivariable Mendelian randomization (MVMR) is a statistical approach using genetic variants as instrumental variables to estimate direct causal effects of multiple exposures on an outcome simultaneously. In univariable MR findings are typically illustrated through plots created using summary data from genome-wide association studies (GWAS), yet analogous plots for MVMR have so far been unavailable due to the multidimensional nature of the analysis. MethodsWe propose a radial formulation of MVMR, and an adapted Galbraith radial plot, which allows for the direct effect of each exposure within an MVMR analysis to be visualised. Radial MVMR plots facilitate the detection of outlier variants, indicating violations of one or more assumptions of MVMR. In addition, the RMVMR R package is presented as accompanying software for implementing the methods described. ResultsWe demonstrate the effectiveness of the radial MVMR approach through simulations and applied analyses, estimating the effect of lipid fractions on coronary heart disease (CHD). We find evidence of a protective effect of high-density lipoprotein (HDL) and a positive effect of low-density lipoprotein (LDL) on CHD, however, the protective effect of HDL appeared to be smaller in magnitude when removing outlying variants. In combination with simulated examples, we highlight how important features of MVMR analyses can be explored using a range of tools incorporated within the RMVMR R package. ConclusionsRadial MVMR effectively visualises causal effect estimates, and provides valuable diagnostic information with respect to the underlying assumptions of MVMR.

13
Correcting for effect modification in the doubly-ranked non-linear Mendelian randomization method

Zhou, A.; Tian, H.; Patel, A.; Mason, A.; Yang, G.; Hypponen, E.; Burgess, S.

2026-01-23 epidemiology 10.64898/2026.01.22.26344640 medRxiv
Top 0.1%
30.7%
Show abstract

The doubly-ranked non-linear Mendelian randomization method can yield biased estimates when instrument strength varies across individuals due to gene-environment (GxE) interactions. We propose a simple strategy to mitigate this bias by modelling GxE interactions and removing the fitted GxE component from the exposure before stratification by the doubly-ranked method. In simulations, the proposed GxE correction strategy eliminated GxE-induced bias with null, linear and non-linear exposure-outcome relationships, and it did not introduce bias even when the effect modifier of the IV-exposure association was a confounder or was correlated with a mediator or collider of the exposure-outcome association. In empirical analyses of serum 25(OH)D, BMI, and LDL-C, falsification tests showed bias in the uncorrected doubly-ranked method. Under the selected panel of effect modifiers, the extent of bias attenuation achieved by GxE correction varied by exposures. GxE correction was most effective for LDL-C, with further support from analyses using negative controls (age at recruitment and sex) and coronary artery disease as a positive control. These findings provide proof of principle evidence that our proposed GxE correction strategy can mitigate GxE-induced bias in practice. Where applicable, we recommend implementing this GxE correction strategy as a sensitivity analysis to assess the robustness of findings from the doubly-ranked method.

14
Bias from heritable confounding in Mendelian randomization studies

Sanderson, E.; Rosoff, D.; Palmer, T.; Tilling, K.; Davey Smith, G.; Hemani, G.

2024-09-06 epidemiology 10.1101/2024.09.05.24312293 medRxiv
Top 0.1%
30.4%
Show abstract

Mendelian randomization (MR) uses genetic variants to estimate the causal effect of an exposure on an outcome in the presence of unmeasured confounding. A key assumption of MR is that the genetic variants used influence the outcome only through the exposure. Violation of this assumption undermines the gene-environment equivalence principle, which posits that modifying the exposure via genetic variation is equivalent to modifying it through environmental factors. With increasing sample sizes in genome-wide association studies genetic instruments with smaller effect sizes are being identified as associated with a trait. Through simulation studies, we demonstrate that such variants may have greater liability to act through confounders of the exposure and outcome in a MR study, biasing effect estimates. This bias acts in the same direction as the confounded associations observed in linear regression, but often with greater magnitude and acts in the same direction across all of the most commonly used MR estimation methods, potentially leading to misleading confidence in the results. We further show that the magnitude of bias escalates as the proportion of genetic instruments associated with confounders increases. Importantly, when potential heritable confounders the genetic variants act through are known and can be instrumented, unbiased causal estimates can be obtained through pre-estimation filtering or by employing multivariable MR and adjusting for the confounder. We illustrate our approach through an application to estimate the effect of C Reactive protein on type 2 diabetes using a hypothesis free approach to identify and remove the effect of potential heritable confounders.

15
Strategies to investigate and mitigate collider bias in genetic and Mendelian randomization studies of disease progression

Mitchell, R. E.; Hartley, A. E.; Walker, V.; Gkatzionis, A.; Yarmolinsky, J.; Bell, J. A.; Chong, A. H. W.; Paternoster, L.; Tilling, K.; Davey Smith, G.

2022-04-22 genetic and genomic medicine 10.1101/2022.04.22.22274166 medRxiv
Top 0.1%
27.0%
Show abstract

Genetic studies of disease progression can be used to identify factors that may influence survival or prognosis, which may differ from factors which influence on disease susceptibility. Studies of disease progression feed directly into therapeutics for disease, whereas studies of incidence inform prevention strategies. However, studies of disease progression are known to be affected by collider (also known as "index event") bias since the disease progression phenotype can only be observed for individuals who have the disease. This applies equally to observational and genetic studies, including genome-wide association studies and Mendelian randomization analyses. In this paper, our aim is to review several statistical methods that can be used to detect and adjust for index event bias in studies of disease progression, and how they apply to genetic and Mendelian Randomization studies using both individual and summary-level data. Methods to detect the presence of index event bias include the use of negative controls, a comparison of associations between risk factors for incidence in individuals with and without the disease, and an inspection of Miami plots. Methods to adjust for the bias include inverse probability weighting (with individual-level data), or Slope-hunter and Dudbridges index event bias adjustment (when only summary-level data are available). We also outline two approaches for sensitivity analysis. We then illustrate how three methods to minimise bias can be used in practice with two applied examples. Our first example investigates the effects of blood lipid traits on mortality from coronary heart disease, whilst our second example investigates genetic associations with breast cancer mortality.

16
Utilising offspring genotype by proxy Mendelian randomization to investigate the causal effect of offspring traits on parental health

Hatton, A.; Brito Nunes, C.; Lawlor, D.; Evans, D.

2025-06-13 genetic and genomic medicine 10.1101/2025.06.11.25329457 medRxiv
Top 0.1%
26.8%
Show abstract

Offspring can exert profound effects on the health of their parents. This is perhaps most apparent during the perinatal period, where the fetus influences processes that alter pre- and post-natal maternal physiology. In theory, it is possible to investigate the causal effect of offspring traits on parental health outcomes using Mendelian randomisation (MR), however, as parental and offspring genotypes are correlated, analyses need to be adjusted for the parents genotype to avoid confounding through the parental genome. Such analyses are difficult to perform at scale because of the paucity of cohorts across the world with large numbers of genotyped maternal- or paternal-offspring dyads and parent-offspring trios. In this manuscript, we explain how the causal effects of offspring traits on parental health outcomes can be investigated using Mendelian randomization (MR) and discuss the challenges in implementing such designs. We introduce the "offspring genotype by proxy" MR framework which can be employed in the absence of offspring genetic information to complement existing approaches in the triangulation of causal inference. The basic idea is to use parental genotypes to proxy the direct effect of their offsprings genotype on their offsprings own exposures. Specifically, we show how it is possible to proxy offspring genotype with paternal genotype when investigating causal effects of offspring traits on maternal health outcomes (and vice versa for paternal outcomes), which minimises the problem of confounding from the relevant parents genotype. We compare our framework to other MR designs that might be used to explore effects of offspring traits on parental health and investigate the consequences of model misspecification and spousal misclassification on statistical power and consistency. Given the increasing availability of datasets like the UK Biobank that (incidentally) include tens of thousands of genome-wide genotyped spousal pairs as well as large population based biobanks with linked health record data for first-degree relatives, we conclude that the offspring genotype by proxy MR approach could augment causal analyses of offspring exposures on their parents outcomes as implementation is not restricted to datasets with parent-offspring genotype information.

17
Developmental origins of exceptional health and survival: A four-generation family cohort study

Keys, M. T.; Pedersen, D. A.; Larsen, P. S.; Kulminski, A.; Feitosa, M.; Wojczcynski, M.; Province, M.; Christensen, K.

2024-05-06 epidemiology 10.1101/2024.05.04.24306872 medRxiv
Top 0.1%
23.0%
Show abstract

Descendants of longevity-enriched sibships demonstrate a broad health and survival advantage throughout the life course. However, little is known about manifestations during very early life. Here we show a pattern of lower risk of adverse early life outcomes in third-generation grandchildren (N = 5637) of Danish longevity-enriched sibships compared to the general population, including infant mortality (Hazard Ratio = 0.53, 95% CI [0.36, 0.77]) and a range of neonatal health indicators. These associations in fourth-generation great-grandchildren (N = 14,908) were strongly attenuated and less consistent (e.g., infant mortality, Hazard Ratio = 0.90, [0.70, 1.17]). These dilatory patterns across successive generations were independent of stable socioeconomic and behavioural advantages (e.g., parental education and maternal smoking), maternal and paternal lines of transmission, as well as secular trends in the background population. Our findings suggest that exceptional health and survival may have early life developmental components and implicate heritable genetic and or epigenetic factors in their transmission. BackgroundPrevious researched has demonstrated potent health and survival advantages across three-generations in longevity-enriched families. However, the survival advantage associated with familial longevity may manifest earlier in life than previously thought. MethodsWe conducted a matched cohort study comparing early health trajectories in third-generation grandchildren (n = 5,637) and fourth-generation great-grandchildren (n = 14,908) of longevity-enriched sibships to demographically matched births (n = 41,090) in Denmark between 1973 and 2018. ResultsLower risk was observed across a range of adverse early life outcomes in the grandchildren, including infant mortality (Hazard Ratio (HR) = 0.53, 95% CI [0.36, 0.77]), preterm birth (Odds Ratio (OR) = 0.82, [0.72, 0.93]), small for gestational age (OR = 0.83, [0.76, 0.90]) and neonatal respiratory disorders (OR = 0.77, [0.67, 0.88]). Relative advantages in parental education and maternal smoking were observed in both generations to a similar degree. However, a much smaller reduction in infant mortality was observed in the great-grandchildren (HR = 0.90, [0.70, 1.17]) and benefits across other outcomes were also less consistent, despite persisting socioeconomic and behavioural advantages. Lastly, maternal, and paternal lines of transmission were equipotent in the transmission of infant survival advantages. ConclusionsDescendants of longevity-enriched sibships exhibit a broad health advantage manifesting as early the perinatal period. However, this effect is strongly diluted over successive generations. Our findings suggest that exceptional health and survival may have early developmental components and implicate heritable genetic and or epigenetic factors in their specific transmission. Key MessagesO_LIPrevious researched has demonstrated potent health and survival advantages across three-generations in longevity-enriched families. However, the survival advantage associated with familial longevity may manifest earlier in life than previously thought. C_LIO_LIIn our study of third and fourth-generation descendants of longevity-enriched sibships, we observed a broad infant health and survival advantage reflected by protection against a diverse range of adverse birth outcomes. C_LIO_LIThese advantages were strongly attenuated between the third and fourth generations, independent of otherwise stable socioeconomic and behavioural parental advantages, as well as maternal and paternal lines of transmission. C_LIO_LIOur findings suggest that familial aggregation of exceptional health and survival may have early life developmental components and triangulate to implicate heritable genetic and or epigenetic factors in their transmission. C_LI

18
Non-linear Mendelian randomization of vitamin D and C-Reactive Protein: an interrogation of methods

Leyden, G.; Hamilton, F.; Sanderson, E.; Davey Smith, G.

2025-09-12 epidemiology 10.1101/2025.09.10.25335520 medRxiv
Top 0.1%
23.0%
Show abstract

Mendelian randomization (MR) is an established epidemiological technique which uses genetic variants to strengthen causal inference regarding modifiable exposures. Non-linear MR is an extension to MR which aims to estimate whether the effect differs across the level of the exposure. Many applications of non-linear MR have focused on Vitamin D as an exposure. Using this technique, the study sample is divided into strata, and separate estimates are calculated in each stratum to estimate causal effects at different levels of the exposure (e.g. Vitamin D). For example, a recent study which applied this method identified an apparent protective effect of Vitamin D on C-reactive protein (CRP) levels for those with poor Vitamin D status. However, recent work has highlighted that the commonly used non-linear MR approaches are susceptible to serious bias, suggesting that further methodological development incorporating extensive simulation and empirical investigation is required. In this paper, we provide a commentary on the sources of bias in non-linear MR methods with a re-examination of the relationship between Vitamin D and CRP as an applied example. We highlight the role of negative controls and non-collider variable-based stratification as potential sensitivity tests to identify potential bias for putative non-linear associations in empirical settings.

19
NeMMo: an improved statistical algorithm for excess all-cause mortality surveillance and monitoring

Lytras, T.; Athanasiadou, M.

2026-08-17 epidemiology 10.64898/2026.08.14.26360477 medRxiv
Top 0.1%
22.7%
Show abstract

Background: Reliable estimation of excess mortality is central to population health surveillance. We introduce NeMMo (New Mortality Model), an evolution of the EuroMOMO model for estimating weekly all-cause expected mortality, and assess its behaviour and performance on empirical data. Methods: NeMMo incorporates population offsets, stratifies observed deaths by age group and models seasonality using a periodic B-spline rather than a Serfling-type sinusoidal function. Baseline weeks are selected by a data-driven procedure minimizing the skewness of the residuals before refitting the model, instead of relying solely on fixed calendar windows. NeMMo enables pooling across age groups, direct age standardization and incorporation of external predictors. We applied NeMMo and EuroMOMO to mortality and population data downloaded from Eurostat for 31 countries from 2015 onwards, excluding the COVID-19 pandemic period from baseline estimation. Results: For most countries NeMMo produced a higher expected mortality baseline that better tracked observed deaths, as well as tighter prediction intervals and higher maximum Z-scores, suggesting improved discrimination of mortality excesses. Z-scores and P-scores during non-pandemic weeks were closer to zero with NeMMo than with EuroMOMO but further elevated during pandemic weeks, providing greater separation between pandemic and non-pandemic mortality. Incorporating population offsets resulted in negative linear trends across all countries, consistent with declining mortality after accounting for demographic changes. The periodic B-spline identified substantial heterogeneity in the shape and timing of seasonal mortality that was not captured by a sinusoidal function. Conclusions: NeMMo provides a flexible and parsimonious framework for all-cause mortality surveillance that improves the established EuroMOMO model and offers theoretical, empirical and practical advantages. It is thus suitable both for detecting short-term spikes and for the long-term, age-adjusted quantification and comparison of mortality excesses that has become increasingly important since the COVID-19 pandemic. The accompanying 'nemmo' package for R facilitates its widespread adoption and application.

20
The Robust Bidirectional Association Between Chronic Lung Disease and Incident Osteoporosis: A Two-Stage Individual Participant Data Meta-Analysis of Three International Longitudinal Cohorts (HRS, SHARE, and ELSA)

Jiang, D.; Bao, J.

2026-03-19 respiratory medicine 10.64898/2026.03.18.26348689 medRxiv
Top 0.1%
22.7%
Show abstract

Abstract Background: The association between chronic lung disease (CLD) and osteoporosis (OP) is well-recognized, but the direction and magnitude of this relationship remain debated, particularly in aging populations. We aimed to quantify the bidirectional association between CLD (including COPD and asthma) and incident OP using a two-stage individual participant data (IPD) meta-analysis of three large longitudinal cohorts. Methods: We harmonized and analyzed individual-level data from the Health and Retirement Study (HRS, USA), the Survey of Health, Ageing and Retirement in Europe (SHARE, Europe), and the English Longitudinal Study of Ageing (ELSA, UK), all comprising adults aged greater than or equal to[≥]50 years. In the first stage, Cox proportional hazards models were fitted separately in each cohort to estimate hazard ratios (HRs) for the forward (CLD[->]OP) and reverse (OP[->]CLD) associations, adjusting for a comprehensive set of confounders (demographics, lifestyle, comorbidities, functional status). In the second stage, cohort-specific log HRs were pooled using fixed-effect meta-analysis. Heterogeneity was assessed with the I-squared statistic. Results: A total of 40,050 participants were included across the three cohorts. The pooled HR for incident OP among individuals with baseline CLD was 1.37 (95% confidence interval [CI] 1.24-1.51), with similar estimates for COPD (HR 1.47, 95% CI 1.27-1.69) and asthma (HR 1.35, 95% CI 1.22-1.50). For the reverse association, baseline OP was associated with increased risk of incident CLD (pooled HR 1.16, 95% CI 1.05-1.29), COPD (HR 1.28, 95% CI 1.11-1.47), and asthma (HR 1.17, 95% CI 1.05-1.30). Heterogeneity was low across all analyses (I2[≤]7.5%). Conclusion: This two-stage IPD meta-analysis provides robust evidence of a bidirectional relationship between CLD and OP in older adults. These findings underscore the need for integrated screening and management of both conditions in aging populations.